Nucleic Acids Research
◐ Oxford University Press (OUP)
Preprints posted in the last 30 days, ranked by how well they match Nucleic Acids Research's content profile, based on 1281 papers previously published here. The average preprint has a 0.79% match score for this journal, so anything above that is already an above-average fit.
Kaufman, P. D.; Liu, H.; Hu, K.; Ferguson, L.; Collins, K.; Zhu, L. J.; Pederson, T.
Show abstract
Various methods have detected miRNA-target interactions via immunoprecipitation of UV-crosslinked Argonaute ribonucleoprotein complexes, followed by intermolecular ligation of bound miRNAs to target strands, forming chimeric RNAs. To date, these methods have relied on conventional viral reverse transcriptases (RTs) to generate cDNAs for sequencing. However, crosslinked RNAs often retain adducts after purification, which can make them poor templates for viral RTs. Here, we adapted OTTR (Ordered Two-Template Relay) techniques to generate cDNAs from Ago2-bound RNAs. OTTR makes use of a modified retroelement-encoded RT, which is strongly processive even on templates with modifications or adducts. We show that this "OTTR-CLASH" method increases the frequency of generating chimeric RNAs compared to previous methods. We also developed an improved bioinformatic pipeline for analysis of these data, and we use this to catalog miRNA-target interactions not previously described in the literature. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=147 HEIGHT=200 SRC="FIGDIR/small/738487v1_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@13bc276org.highwire.dtl.DTLVardef@5beb41org.highwire.dtl.DTLVardef@b204e5org.highwire.dtl.DTLVardef@15f747d_HPS_FORMAT_FIGEXP M_FIG C_FIG
Conklin, D.; Lee, J.-A.; Palazzolo, M.; Dubinett, S. M.; Lee, J. M.
Show abstract
Targeted knock-in technologies have enabled precise insertion of reporters, affinity tags, degrons, and other functional payloads into endogenous genomic loci. Over the past decade, a diverse collection of genome engineering strategies has emerged, including approaches based on homology-directed repair (HDR), microhomology-mediated end joining (MMEJ), homology-mediated end joining (HMEJ), and related methodologies. While these advances have greatly expanded the capabilities of endogenous genome engineering, they have also increased the complexity of donor design, assembly, and validation. Here, we describe FORGE-KI (Functional Oncology Research Genetic Engineering - Knock in), a pathway-matched design workflow for endogenous knock-in engineering that aligns the assembly strategy with the underlying repair mechanism. For large-cargo insertions, we use a modular five-component framework that separates gene-specific targeting arms from reusable functional modules, allowing rapid assembly of HDR donor constructs targeting AHR, IRF1, and FOSL1 from a shared reagent collection. For MMEJ/PITCh applications, where short targeting elements permit rapid fabrication, we developed a streamlined one-step pipeline in which the entire donor and selection payload is synthesized as a single continuous fragment for direct cloning, compressing the design-to-reagent cycle time. This MMEJ workflow is paired with a dual-promoter nuclease vector (pForge-KI-MMEJ-Cas9-DualGuide) that drives the PITCh-release and locus-specific guides from distinct promoters, a design intended to reduce the repeated-promoter instability associated with some dual-guide vectors. We also established a standardized workflow for donor assembly, generation of knock-in cell populations, molecular validation, and selectable-cassette removal, and we demonstrate it by generating a functional, selection-marker-free, cytokine-inducible IRF1 HDR reporter line and an inducible IRF1 PITCh/MMEJ reporter pool with confirmed junction enrichment. In parallel, we developed forgeKI, an R package that automates C-terminal reporter knock-in design across both HDR and PITCh/MMEJ repair pathways, including guide selection, target-biology validation, targeting-arm design, domestication, donor-assembly planning, and generation of synthesis-ready constructs. Together, the reagents and software provide a practical system for endogenous knock-in engineering that supports multiple payloads, selection strategies, and repair pathways within a shared donor organization. Rather than replacing existing knock-in technologies, this framework provides a modular foundation for incorporating, extending, and automating the published knock-in methods.
Walters-Freke, C.; Hoshika, S.; Perry, A.; Benner, S.; Dobson, R.; Tillett, Z.; Richards, N.; Williamson, A.
Show abstract
Artificially Expanded Genetic Information Systems (AEGIS) increase the information content of nucleic acids by including new nucleobase pairings that are orthogonal to those of canonical Watson-Crick nucleobases. DNA ligases do not form direct interactions with the nucleobases during catalytic turnover, suggesting that these enzymes should efficiently and faithfully join double-stranded AEGIS substrates. Here we report the systematic investigation into the validity of this hypothesis for structurally-diverse DNA ligases employing substrates built from the eight nucleotide hachimoji genetic alphabet, where orthogonality is achieved by rearranging the hydrogen bonding patterns seen in canonical Watson-Crick pairs. We find that single, or multiple, non-canonical bases are well tolerated at the 5 prime-end of the nick. However, tracts of consecutive non-canonical bases at the 3 prime-end of the break significantly decrease ligation efficiency or abolish it altogether. Possible reasons for this apparent bias against non-canonical nucleobases could include incompatibility in electrostatic interactions between the ligase active site and the non-canonical substrates or altered conformational preferences and/or dynamics in key catalytic intermediates. We also observe single hachimoji mismatches are ligated more frequently than mis paired canonical bases, potentially due to promiscuous pairing of tautomeric forms of the non-canonical bases.
Gravel, C. M.; Berry, K. E.
Show abstract
The bacterial three-hybrid (B3H) assay is a powerful genetic tool for detecting interactions between RNA and RNA-binding proteins (RBPs) and assessing the consequences of RBP mutations. This transcription-based system connects the strength of an RNA-protein interaction to the expression of a lacZ reporter gene in Escherichia coli cells. This in vivo approach allows researchers to dissect RNA-protein interactions within a cellular environment, bypassing the need for biochemical purification of RNAs or proteins. This chapter details a three-day protocol for generating quantitative B3H data. Since a significant challenge in B3H assays is RNA misfolding, we describe a recently optimized set of B3H constructs that mitigates this issue by isolating bait RNAs as discrete folding units.
Nie, L.
Show abstract
Compact tissue-specific promoters are highly desirable for gene therapy because viral vectors possess limited packaging capacity. However, existing promoter engineering strategies rely primarily on rational design or de novo sequence generation and lack efficient approaches for compressing long native promoters while preserving regulatory specificity. Although genome foundation models have substantially improved sequence-to-function prediction, they have not been effectively translated into computational platforms for promoter engineering. Here, we present VirEvo, a computational promoter engineering framework that integrates a virtual dual-luciferase assay (VirDLA), genome-foundation-model-guided genetic evolution, and an orthogonal Pan-Tissue Consistency Filter (PTCF). VirDLA introduces an internal-reference normalization strategy inspired by dual-luciferase reporter assays, enabling relative comparison of promoter activity across tissues without retraining AlphaGenome. Guided by these normalized activity scores, VirEvo iteratively optimizes promoter selectivity, off-target activity, and sequence length. Using the human p16INK4a promoter as a proof of concept, VirEvo evolved a compact synthetic promoter, SRP2M, of only 398 bp, representing an 85.9% reduction in sequence length. Experimental validation using dual-luciferase reporter assays in senescent IMR90 fibroblasts demonstrated that SRP2M retained 77% of wild-type senescence selectivity while reducing basal leakage to 52% of the wild-type level. Together, these results demonstrate the feasibility of genome-foundation-model-guided promoter engineering. VirEvo provides a generalizable framework for designing compact tissue-specific regulatory elements and extends the application of genome foundation models from functional prediction to synthetic regulatory engineering.
Dysart, M. J.; Fang, L.; Karinje, L. K.; Chappell, J.; Stadler, L. B.; Silberg, J. J.
Show abstract
TEXT ABSTRACTCatalytic-RNA (cat-RNA) expressed from mobile DNA can record cellular events, such as the uptake of plasmids via horizontal gene transfer, by splicing a barcode onto 16S ribosomal RNA (rRNA) - a system termed RNA addressable modification (RAM). However, scaling RAM to record multiple simultaneous biological events requires large numbers of orthogonal cat-RNA whose signals reflect the biological features under investigation rather than variability arising from the barcode sequence. Here, we explore how to design orthogonal cat-RNA to record information about multiple plasmid-encoded traits in parallel. We show that cat-RNA having tRNA-derived barcodes with sequence variation in the anticodon stem-loop present greater signal consistency within Escherichia coli than mRNA-derived barcodes. When orthogonal cat-RNA designs harboring tRNA-derived barcodes were evaluated in Vibrio natriegens and Pseudomonas putida, increased variance was observed compared with Escherichia coli. Nevertheless, the signal consistency was sufficient to use these orthogonal cat-RNAs to report on the relative activities of four promoters and two origins of replication by sequencing barcoded-rRNA derived from the three organisms. These results show how RAM can be multiplexed to report on mobile DNA features in microbial communities and illustrate the importance of accounting for variability in RNA outputs when designing and interpreting multiplexed RNA barcoding data. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=88 SRC="FIGDIR/small/738544v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@406ebaorg.highwire.dtl.DTLVardef@259751org.highwire.dtl.DTLVardef@1f1512corg.highwire.dtl.DTLVardef@8384b_HPS_FORMAT_FIGEXP M_FIG C_FIG
Rusinek, W.; Dorawa, S.; Kaczorowski, T.
Show abstract
Thermostable DNA polymerases are indispensable tools in molecular biology, yet enzymes from the most extreme hyperthermophiles remain largely uncharacterized. Here, we report the biochemical and structural characterization of a family B DNA polymerase from Pyrolobus fumarii A1 (Pyrfu pol), one of the most thermoresistant archaea described to date. The enzyme was efficiently overproduced in E. coli Rosetta 2(DE3)[pLysS] and purified to homogeneity using a two-step protocol that combined heat treatment with immobilized metal affinity chromatography (IMAC). Bioinformatic analysis confirmed the canonical family B architecture, while AlphaFold-based structural modeling and comparative analysis with mesophilic RB69 DNA polymerase revealed a well-conserved structural core alongside thermoadaptive features. Radiolabel incorporation assays demonstrated enzymatic activity over a broad ionic strength range and an absolute requirement for Mg ions. PCR-based optimization confirmed these findings and revealed broad pH tolerance (6.5-11.0). Notably, Tris inhibited radiolabel-based assays (pH 7.0) yet proved essential for efficient PCR amplification (pH 8.5), suggesting a context-dependent role of buffer composition in polymerase activity. Processivity assays confirmed amplification of DNA fragments up to approximately 8,000 bp. Replication fidelity, assessed by the lacZ-based assay, showed a 2.9-fold improvement over Taq polymerase. Urea-nanoDSF yielded an exceptional melting temperature of 105.9 {+/-} 0.08 {degrees}C. Pyrfu pol also demonstrated tolerance to common PCR inhibitors, highlighting its potential utility in molecular biology applications.
Hazra, A. B.; Kalita, D. B.; Bhattacharyya, A.; Gupte, V.; Venugopal, V.; Pattathil, A.
Show abstract
S-adenosyl-L-methionine (SAM), an essential cofactor in all forms of life, is synthesized by the enzyme methionine adenosyltransferase (MAT) from methionine and ATP. The adenine moiety in SAM appears to have no direct function in catalysis, and some MAT homologs can utilize natural nucleotide triphosphates in vitro, producing the corresponding SAM nucleobase analogues. However, the molecular determinants of nucleotide choice of the MAT enzyme and the cellular significance of the nucleobase in SAM are unclear. In this study, using structure- and bioinformatics-guided mutagenesis, we identify a flexible active-site loop as a major determinant of nucleotide specificity in MAT. Loop mutations and loop swaps convert ATP-selective Escherichia coli MAT into variants that accept GTP, CTP, and UTP, enabling enzymatic synthesis and purification of S-guanosyl-, S-cytosyl-, and S-uracyl-L-methionine. Further, we show that these analogues partially rescue the growth of an E. coli SAM auxotroph under SAM-limited growth conditions. Biochemical assays show that the analogues bind the tested SAM-utilizing enzymes; they serve as substrates for E. coli SAM decarboxylase but do not support detectable methyl transfer by E. coli DNA adenine methyltransferase. These results establish the flexible loop as a gatekeeper of MAT nucleotide specificity and show that this loop can be engineered to produce SAM analogues which can selectively participate in downstream cellular metabolism. Graphical Abstract/ Table of contents only O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=108 SRC="FIGDIR/small/737877v1_ufig1.gif" ALT="Figure 1"> View larger version (46K): org.highwire.dtl.DTLVardef@eea079org.highwire.dtl.DTLVardef@698836org.highwire.dtl.DTLVardef@6da3e7org.highwire.dtl.DTLVardef@239910_HPS_FORMAT_FIGEXP M_FIG C_FIG
Mansour, A.; Sarigul, I.; Tenson, T.; Maivali, U.
Show abstract
The Escherichia coli protein YbeX/CorC is encoded in the same operon as the ribosome biogenesis factor YbeY, and its deletion leads to accumulation of 17S pre-rRNA and degradation intermediates of 16S rRNA under magnesium limitation. To further investigate the ybeX deletion phenotype, we used rRNA fluorescence in situ hybridization coupled with flow cytometry (rRNA-FISH-flow) to quantify 16S rRNA, 23S rRNA, and 17S pre-rRNA levels at single-cell resolution in{Delta} ybeX and{Delta} ybeY strains.{Delta} ybeX cells grown under limiting Mg2+ develop striking cell-to-cell heterogeneity in 17S pre-rRNA content during the transition to stationary phase, with up to 25-fold differences between individual cells. Upon regrowth from the stationary phase,{Delta} ybeX cultures display a bimodal distribution of 17S pre-rRNA, revealing two distinct subpopulations -- one retaining high levels of unprocessed pre-rRNA and the other with low levels -- whose relative proportions shift over time, until visible growth resumes. The stoichiometry between mature 16S and 23S rRNAs remains tight in both strains, indicating that the heterogeneity is specific to pre-rRNA processing, rather than a general disruption of ribosome homeostasis. The{Delta} ybeY mutant accumulates 17S pre-rRNA more uniformly across cells and primarily during exponential growth in rich medium, consistent with its direct role in 16S rRNA maturation. These single-cell data suggest that YbeX and YbeY affect ribosomal RNA metabolism through distinct mechanisms and that the extended lag phase of{Delta} ybeX is caused by a heterogeneous clearing of pre-ribosomal intermediates in individual cells.
Kuryavyi, V. V.
Show abstract
Abstract The universe of possible nucleotide sequences expands combinatorially with sequence length, vastly exceeding the fraction sampled by real genomes. Yet genomic sequences exhibit reproducible compositional symmetries and recurrent structural motifs, indicating that biological sequence space is shaped by strong organizing constraints. Here, we introduce an explicit framework for constructing and visualizing the complete sequence universe using the Newtonian polynomial for a four-letter alphabet, and for identifying biologically relevant subsets through the application of fundamental filters. Three filters of biological relevance are formulated: (i) the constraint that DNA predominantly exists as an antiparallel-stranded double helix, (ii) the second Chargaff parity rule, which enforces approximate strand symmetry in single-stranded sequence composition, and (iii) genome shadows, reflecting the imprint of concerted sequence changes. Successive application of these filters dramatically reduces the accessible sequence space and reveals distinct symmetry classes. Among these, mirror-symmetric sequences occupy a privileged position because they are invariant under strand reversal and therefore compatible with both antiparallel and parallel strand orientations. This dual compatibility enables such sequences to bridge otherwise disjoint structural subspaces of DNA. G-rich members of this class are shown to have a strong propensity to form G-quadruplex architectures that incorporate parallel-stranded domains while remaining compatible with duplex DNA. We propose that this structural versatility provides a mechanistic basis for the recurrent association of G-rich mirror-symmetric sequences with recombination hotspots and genome rearrangements. Together, these results establish a symmetry-based framework for understanding how combinatorial sequence space is filtered into biologically functional DNA motifs.
Serdakov, M. D.; Bohdan, D. R.; Nikolaev, G. I.; Bujnicki, J. M.; Baulin, E. F.
Show abstract
Non-coding RNAs play diverse roles in a wide range of cellular processes, with their spatial structure being pivotal to their function. RNA secondary structure is a key determinant of its overall fold. Given the scarcity of experimentally determined RNA 3D structures, understanding secondary structure is vital for discerning RNA function. Currently, there is no universally effective solution for de novo RNA secondary structure prediction. Existing methods are becoming increasingly complex without marked improvements in accuracy and often overlook critical features such as pseudoknots and alternative folds. Here, we introduce SQUARNA, a new approach to de novo RNA secondary structure prediction that is suitable for both individual RNA analysis and large-scale structural searches. SQUARNA revisits the concept of base pair maximization and develops it into a stem maximization idea coupled with the widely used free energy minimization (MFE) framework. SQUARNA can predict alternative structures and handle pseudoknots of arbitrary complexity. Benchmarking shows that SQUARNA outperforms existing methods, including deep learning models, in both single-sequence and alignment-based RNA secondary structure prediction. SQUARNA seamlessly integrates sequence and alignment information with experimental data, such as residue reactivities obtained by chemical probing, as well as other structural restraints, including automated searches for Rfam database templates, G-quadruplex patterns, and protein-binding motifs. SQUARNA is available as a standalone tool at https://github.com/febos/SQUARNA and as a web server at https://larnal.imol.institute.
Pronk, B.; Makrodimitris, S.; Wilting, S.; Reinders, M.
Show abstract
MotivationAccurate discrimination between healthy individuals and patients with cancer using minimally invasive liquid biopsies could improve cancer diagnosis and monitoring. Circulating cell-free DNA (cfDNA) is a promising biomarker, since fragmentation patterns reflect chromatin organization and have been used to interrogate regulatory regions such as transcription start sites (TSSs). Classification approaches typically rely on hypothesis-driven selection of genomic regions based on literature or external tissue data. Therefore, they assume that tumor-derived cfDNA constitutes the dominant diagnostic signal, potentially overlooking a systemic, genome-wide shift in the cfDNA pool. ResultsWe present a data-driven framework that identifies discriminative genomic loci directly from cfDNA whole-genome sequencing data. Using fragmentomic features captured at TSSs within a nested cross-validation framework, the model outperforms ichorCNA and hypothesis-driven baselines in distinguishing healthy from colorectal and breast cancer samples (AUROC 0.95 {+/-} 0.039). Performance was maintained in a pan-cancer setting across seven malignancies (AUROC 0.946 {+/-} 0.032) and generalized to previously unseen cancer types within the same cohorts (AUROC 0.934 {+/-} 0.006). While validation in an independent external cohort showed a performance gap (AUROC 0.694), the data-driven model was consistently competitive with baseline methods. These results indicate that robust cancer detection is enabled by integrating distributed genome-wide fragmentation patterns rather than restricting analysis to predefined regions. Availability and implementationScripts to reproduce the results are available at https://github.com/brmprnk/comp/ Contacti.b.pronk@tudelft.nl Supplementary informationavailable at NAR Online.
Christopoulou, N.; Dương, N. H.; Arede-Rei, P.; Torrens, G.; Blandenet, M.; Cava, F.; Granneman, S.
Show abstract
Analysis of RNA-binding proteome data from different bacterial species revealed many cell wall metabolic enzymes cross-linking to RNA in vivo, hinting that these proteins directly bind RNA. Surprisingly, penicillin-binding proteins (PBPs) were also abundantly identified as putative RNA-binding proteins. The cell surface localisation properties of many of these proteins therefore beg the question at what stage of their cellular life cycle these proteins interact with RNA and what the functional significance is. Here, we characterised the RNA-binding activity of PBP2a, the alternative transpeptidase that confers {beta}-lactam resistance in MRSA. Using in vivo RNA-binding assays, we show that PBP2a interacts with hundreds of transcripts without apparent sequence specificity. Computational analyses identified a possible RNA-binding cleft in PBP2a proximal to its active site. Mutation of only two predicted positively charged residues located in this cleft substantially reduced cross-linking in vivo, implying that RNA recognition is largely dictated by RNA backbone interactions. While PBP2a does not regulate RNA steady-state levels, RNA-binding appears important for proper protein function: an RNA-binding deficient mutant exhibits reduced oxacillin resistance. These findings establish PBP2a as an RNA-binding protein in vivo and provide a framework to investigate how this non-canonical interaction may relate to cell wall biogenesis and {beta}-lactam resistance.
Lari, A.; Shah, S. B.; Batarseh, S.; Nagorsen, M.; Glaunsinger, B. A.
Show abstract
Cells must be primed to rapidly induce inflammatory gene expression upon infection while also tuning the level of induction to avoid immunopathology. Here, we identify RNA polymerase III (Pol III), best known for transcribing noncoding RNAs, as a dual-function regulator of RNA polymerase II (Pol II)-dependent inflammatory gene expression. Pol III is selectively enriched at promoters of innate immune, pro-inflammatory, and stress-response genes, where it maintains chromatin accessibility and supports basal transcription. Upon infection with murine gammaher-pesvirus 68 (MHV68), Pol III redistributes from these promoters to retrotransposon loci, coinciding with enhanced expression of inflammatory genes. Depletion of the Pol III transcription factor Brf1 further amplifies inflammatory transcription during infection with MHV68, herpes simplex virus-1, and influenza A virus. Genes restrained by Pol III have TATA-box-enriched promoters and are functionally dependent on TATA-binding protein (TBP), suggesting that Pol III modulates inflammatory gene expression by competing with Pol II for shared transcriptional machinery. Thus, Pol III is a chromatin licensor in uninfected cells and an inflammation rheostat during viral infection. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=178 SRC="FIGDIR/small/738346v1_ufig1.gif" ALT="Figure 1"> View larger version (43K): org.highwire.dtl.DTLVardef@d99325org.highwire.dtl.DTLVardef@4b781dorg.highwire.dtl.DTLVardef@bad1b6org.highwire.dtl.DTLVardef@11e41a6_HPS_FORMAT_FIGEXP M_FIG C_FIG
Arya, A.; Datta, B.
Show abstract
Symmetry elements in nucleic acids are most strongly correlated with sites of biological function; however, their relevance to non-canonical structures remains underexplored. In this study, we demonstrate the presence and significance of trinucleotide symmetry elements within G-quadruplex (G4) motifs. Our central hypothesis is that the intra-strand mirror symmetry of trinucleotides has been evolutionarily selected to facilitate G4 formation builds on the established sequence-structure association of G-quadruplexes and the natural symmetry law governing nucleotide insertion during genome evolution. Using a conserved G4 motif in the first exon of the MTOR gene as a model, we showed remarkable trinucleotide symmetry preservation across primates and broader mammals, with functional G4 regions displaying locally elevated symmetry relative to the codon-biased exonic background. Analysis of experimentally validated oncogenic G4s, including c-MYC, BCL2, VEGF, and KRAS, revealed that mirror and reverse complement symmetries converge around biologically important G4s. To quantify this feature, we formulated two complementary descriptors: the mirror symmetry index (MSI) and its non-palindromic variant (nMSI). Across 14 oncogene-promoter wild-type G4s, the majority scored MSI [≥] 0.80 (mean 0.884), with only the loop-rich ATG7, BCR, and MDM2 motifs falling below this value, and the KRAS promoter G4 reached individual significance against its mononucleotide-preserving null distribution (p = 0.042). Most decisively, each wild-type G4 scored higher on MSI than its experimentally confirmed G4-abolished mutant in 12 of 14 paired comparisons (sign test, p = 0.0065; mean {Delta}MSI = +0.089, mean {Delta}nMSI = +0.192); the two reversals (BCL2 and HIF-1) are attributable to scrambled mutant controls that introduce more balanced trinucleotide compositions rather than to failure of the index. The directional trend was reproduced across three independently published datasets, with nMSI [≥] 0.50 separating G4-forming from non-G4 sequences at 77.8% sensitivity and 100% specificity, although the collective per-sequence signal from mononucleotide-preserving shuffles remained a non-significant trend (Stouffer combined Z = 1.197, p = 0.116). This first report of trinucleotide symmetry in G4 motifs posits that coordinated nucleotide insertion and quadruplet maintenance act as an evolutionary forcing mechanism that pre-organizes single strands for G4 folding.
Gromak, D.; Shaytan, A. K.; Herbert, A.; Poptsova, M.
Show abstract
The p150 isoform of the double-stranded RNA editing enzyme ADAR1 binds Z-DNA and Z-RNA through the conserved winged helix-turn-helix Z domain. Here, we describe an inverse computational design strategy to map protein interactors of Z. We used RFdiffusion and ProteinMPNN to generate [~]10,000 synthetic binders optimized for the Z recognition surface, then used their sequences as structural templates for BLASTp searches against the human proteome. Multi-stage screening of [~]1,200 candidate regions from 298 proteins via ColabFold pDockQ identified 79 candidates for high-resolution AlphaFold3 modeling, which revealed the m6A reader YTHDC1 as the top-ranked interactor. AlphaFold3 predicts that a glutamate-rich poly-E disordered region of YTHDC1 (residues 199-254) docks into the basic recognition pocket of Z through a charge-complementary mechanism that mimics the phosphate backbone of Z-RNA. Microsecond molecular dynamics simulations confirmed stability of the binary ADAR1p150-YTHDC1 complex, with the Z-poly-E interface maintaining RMSD below 3 [A] throughout. Ternary complex simulations showed that dsRNA acts as a co-anchoring scaffold stabilizing simultaneous engagement of both proteins in a catalytically dormant conformation. YTHDC1 localizes to transcription-associated YT bodies where nascent RNAs undergo m6A modification and negative supercoiling promotes Z-DNA formation, suggesting that YTHDC1 recruits ADAR1p150 to promote editing of intron-containing substrates prior to splicing. Short StatementADAR1 can edit RNAs after they are made, changing the message they carry and removing double-stranded RNAs that can activate inflammatory responses. Using AI-driven protein design, we discovered that ADAR1 physically interacts with YTHDC1, a protein that recognizes newly made self-RNAs. The interaction allows ADAR to edit RNAs as they are made. This unexpected connection between two fundamental RNA modification systems opens new avenues for understanding and potentially targeting autoimmune disorders and cancers driven by the mis-editing of RNA transcripts.
Barmada, M. I.; Hanna, A.; Bair, C. R.; McGinity, E. N.; Zelinskaya, N.; Dey, D.; Conn, G. L.
Show abstract
Bacterial ribosomal RNA (rRNA) methylations are important for accurate translation. Four distinct methylations incorporated by RsmE, RsmF, and RsmH/ RsmI form a cluster of three modified 16S rRNA nucleotides (m3U1498, m5C1407, and m4Cm1402) surrounding the decoding center of the 30S subunit. Given their common substrate requirement of a late-stage intermediate 30S subunit, these enzymes likely act contemporaneously during subunit biogenesis, but whether there exists a required modification order is unknown. Here, using hypomethylated 30S subunits obtained from a collection of rsmH/I/E/F-deleted Escherichia coli strains, we identify RsmF activity to be highly dependent on prior modification of h44 both in vitro and in E. coli. RsmF activity on hypomethylated 30S subunits could be partially rescued by prior in vitro methylation using RsmE and RsmH, indicating that incorporation of these methyl groups directly shapes h44 for recognition by RsmF. RNA structure probing using SHAPE-MaP and molecular dynamics simulations reveal specific alterations in 16S rRNA structure and dynamics in the absence of the m4C1402 (RsmH) and m3U1498 (RsmE) modifications that likely restrict RsmF action. These studies thus uncover a previously unappreciated "order of operations" for 16S rRNA modification during ribosome biogenesis with important implications for studies on the collective functions of these modifications.
Srinivasan, S.; Chande, A.
Show abstract
Post-transcriptional chemical modifications of RNA, collectively termed the epitranscriptome, have emerged as critical regulatory layers governing viral replication, pathogenicity, and host-virus interactions. Despite the rapid accumulation of experimental data on viral RNA modifications, no dedicated, freely accessible resource existed for systematically cataloguing these sites across diverse viral species. Here we present ViralEpiBase, a manually curated database of epitranscriptomic modification sites identified in viral RNA genomes and virus-encoded transcripts at single-nucleotide resolution. ViralEpiBase currently integrates seven chemically distinct RNA modification types: N6-methyladenosine (m6A), N1-methyladenosine (m1A), pseudouridine ({Psi}), 5-methylcytosine (m5C), 2'-O-methylation (2'OMe), inosine and N4-acetylcytidine (ac4C); across 12 viral species encompassing both DNA and RNA viruses of clinical and biological significance. Each entry is linked to its primary literature source or deposited dataset and is retrievable by modification type, genomic coordinates, or viral taxonomy. The database is freely accessible through an intuitive web interface and is updated continuously as new experimental evidence becomes available. ViralEpiBase thus provides the first unified platform dedicated exclusively to viral epitranscriptomics and is designed to facilitate mechanistic investigation of RNA modification functions in viral biology.
Prochownik, E. V.; Henchy, C. M.; Wang, H.
Show abstract
MYC oncoprotein binding at promoters and enhancers influences RNA polymerase II (RNAPII)-driven gene expression. Numerous genes also bind MYC near their transcriptional end sites (TESs). This often allows direct promoter-TES contact via looping and further regulates total and 'read-through' transcription that extends beyond standard termination sites. We aimed here to better clarify the rules governing TES associated MYC and/or RNAPII binding cross-talk in human and murine cells. Using ChIPseq and RNAseq datasets from the ENCODE portal and elsewhere, MYC and RNAPII binding profiles were found to differ around TESs and transcriptional start sites (TSSs). Variations in E box flanking sequences likely accounted for the somewhat lower affinities of MYC for TES-associated sites. Motifs for numerous other transcription factors were also observed to cluster non-randomly and in close proximity to MYC and RNAPII binding site peak summits. On average, genes with TES-proximal MYC or RNAPII sites were more highly expressed than those without, although co-binding tended to be suppressive. Both normal and neoplastic proliferative stimuli altered the MYC and RNAPII binding patterns of many genes, indicating that 'category switching' was common, subject to disparate external signals and often reversible. Functionally related gene sets with high levels of read-through transcription were uniformly marked by significant amounts of TES-associated MYC and/or RNAPII binding. These findings indicate that, both independently and together, MYC and RNAPII binding near TESs dynamically impact total and read-through transcription while also coordinating the expression of many common purpose gene sets.
Lujan-Rodriguez, C.; Popoloski, M. A.; Couturier, L. E.; Richa, J. J.; Talluto, J. M.; Lapine, M. E.; Roche, M.; Edouard, S. J.; Pavan, V.; Kuehner, J. N.
Show abstract
Premature termination of transcription (PTT), also known as attenuation, is a conserved gene regulatory mechanism that operates across all domains of life and in viruses. Attenuation enables rapid cellular responses to environmental and metabolic changes and fine-tunes expression of biosynthetic genes. In Saccharomyces cerevisiae, attenuation of RNA Polymerase II (Pol II) transcription was first linked to the Nrd1-Nab3-Sen1 (NNS) termination pathway for non-coding RNAs, and the mRNA 3-end processing factor Hrp1 has been implicated more recently. Substitutions in Hrp1 RNA Recognition Motifs (RRMs) cause attenuator readthrough and reduce RNA-binding affinity in vitro, but direct evidence for Hrp1 functioning at attenuators in vivo remains limited. Here, we characterized 5-end RNA terminator elements from several genes, including RAD3, SNG1, MNR2, and CPR8. Readthrough mutations clustered in AU-rich regions resembling polyadenylation site (pA) efficiency elements, consistent with Hrp1 binding targets. Amino acid substitutions of Hrp1 RRM residue F162 revealed a general requirement for aromaticity in RNA recognition that varied to some degree by gene context. To test Hrp1-RNA interactions independent of other yeast factors, we adapted a bacterial 3-hybrid (B3H) assay. Hrp1 interacted with RNA derived from the GAL7 3-end pA site and 5-end terminator regions of RAD3, MNR2, and CPR8. Mutations in AU-rich RNA regions that disrupted Pol II attenuation in yeast generally impaired B3H interactions. However, some Hrp1 mutants (M191T, I270T, D271G, M275V, T280I) retained binding to CPR8 terminator RNA, suggesting their defects require additional yeast components. These results demonstrate that Hrp1 is sufficient to bind multiple UA-rich attenuator RNAs in vivo, expanding Hrp1 function to include early transcription events.